Back

Modern Pathology

Elsevier BV

All preprints, ranked by how well they match Modern Pathology's content profile, based on 22 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Harnessing Pathology Foundation Models to Accelerate Lymphoma Diagnosis Through Automated Immunohistochemistry Triage

Zhu, M.; Li, A.; Safa, I.; Galera, P.; Hazoglou, M.; Vanderbilt, C.; Kamali, A.; Goldgof, G.; Veeraraghavan, H.; Jiang, J.; Ardon, O.; Geneslaw, L.; Dogan, A.

2026-08-12 pathology 10.64898/2026.08.11.26360085 medRxiv
Top 0.1%
53.6%
Show abstract

Pathologic diagnoses of hematopoietic diseases require immunohistochemistry (IHC) stains selected by pathologists upon preview of H&E-stained slides. This multi-step workflow can delay diagnostic turnaround time by days. Hence, we developed the Hematopathology Automatic Triaging System (HATS), which automates IHC panel ordering directly from H&E whole-slide images using pretrained pathology foundation model representations combined with attention-based multiple-instance learning. After the most comprehensive evaluation of pathology foundation models for hematologic malignancy classification to date, encompassing seven publicly available models, we trained HATS on 4,996 whole-slide images from 1,607 patients spanning the ten most common lymphoma diagnostic categories. HATS achieves 84% case-level subtype classification accuracy (0.962 ROC-AUC), translating to 92% IHC panel ordering accuracy. In a blinded reader study, HATS outperforms practicing pathologists at predicting lymphoma subtypes from morphology alone (85% vs 65%). In an independent real-world validation of 230 clinical cases, after directing 7 cases with scant tissue for manual review, HATS-ordered IHC panels were sufficient for diagnosis in 72.6% of cases. By automating the triaging step while preserving full pathologist oversight, HATS offers a safe and practical entry point for clinical AI adoption in pathology.

2
An Automated, Pathologist-free Gleason Grade Stratifies Disease-free Interval Comparably to Expert Grading from a Single Out-of-distribution Slide

Ebbert, J. L.; Szymanski, J.; Perry, A.; Della Corte, D.

2026-06-24 pathology 10.64898/2026.06.22.26356247 medRxiv
Top 0.1%
52.5%
Show abstract

Automated Gleason grading now matches expert pathologists on the cohorts where systems are developed and tuned, but deployment-relevant gaps remain: whether an automated grade, applied without site-specific tuning or pathologist oversight, stratifies outcome comparably to expert grading on slides from unseen institutions and in cross-specimen applications. We tested this for disease-free interval (DFI), a curated recurrence endpoint. A production gland-level prostate diagnostic (PathTools Prostate v11.0) was applied frozen and uncalibrated to 298 diagnostic whole-slide images from 274 TCGA-PRAD radical-prostatectomy patients, a cohort outside its development distribution and needle-core-biopsy training data, contributed by 25 source sites under heterogeneous digitization; tissue was detected automatically with no expert region annotation. From the output we derived an ISUP grade group and continuous high-grade content, and evaluated each grade as a standalone predictor of DFI (24 events) by Harrell's c-index with 95% bootstrap confidence intervals, a paired between-method bootstrap, and Kaplan-Meier curves with the log-rank test. The automated grade reproduced the clinical grade group at quadratic-weighted kappa = 0.62 (95% CI 0.53-0.70; 48% exact, 86% within one group), within the expert inter-observer range. As the sole predictor it stratified recurrence (log-rank p = 0.022; c-index 0.69, 95% CI 0.58-0.79), and the continuous high-grade fraction was robustly prognostic (hazard ratio 1.37 per SD, p = 0.029; c-index 0.71, 0.61-0.81). Standalone discrimination was not statistically separable from the clinical grade (c-index 0.78, 0.69-0.86; paired {triangleup} c-index spanning zero), and in a joint model the automated grade added nothing beyond it, consistent with both measuring a shared morphological axis. From a single out-of-distribution slide with no pathologist oversight, the automated grade provides standalone recurrence stratification not statistically separable from whole-gland expert grading, demonstrating robust generalizability beyond training data; reported as a continuous high-grade fraction, it offers reproducible, expert-free, grade-equivalent risk stratification for harmonizing large archival or genomically-profiled cohorts.

3
iQC: machine-learning-driven prediction of surgical procedure uncovers systematic confounds of cancer whole slide images in specific medical centers

Schaumberg, A. J.; Lewis, M. S.; Nazarian, R.; Wadhwa, A.; Kane, N.; Turner, G.; Karnam, P.; Devineni, P.; Wolfe, N.; Kintner, R.; Rettig, M. B.; Knudsen, B. S.; Garraway, I. P.; Pyarajan, S.

2023-12-13 pathology 10.1101/2023.09.19.23295798 medRxiv
Top 0.1%
52.1%
Show abstract

ProblemThe past decades have yielded an explosion of research using artificial intelligence for cancer detection and diagnosis in the field of computational pathology. Yet, an often unspoken assumption of this research is that a glass microscopy slide faithfully represents the underlying disease. Here we show systematic failure modes may dominate the slides digitized from a given medical center, such that neither the whole slide images nor the glass slides are suitable for rendering a diagnosis. MethodsWe quantitatively define high quality data as a set of whole slide images where the type of surgery the patient received may be accurately predicted by an automated system such as ours, called "iQC". We find iQC accurately distinguished biopsies from nonbiopsies, e.g. prostatectomies or transurethral resections (TURPs, a.k.a. prostate chips), only when the data qualitatively appeared to be high quality, e.g. vibrant histopathology stains and minimal artifacts. Crucially, prostate needle biopsies appear as thin strands of tissue, whereas prostatectomies and TURPs appear as larger rectangular blocks of tissue. Therefore, when the data are of high quality, iQC (i) accurately classifies pixels as tissue, (ii) accurately generates statistics that describe the distribution of tissue in a slide, and (iii)accurately predicts surgical procedure from said statistics. We additionally compare our "iQC" to "HistoQC", both in terms of how many slides are excluded and how much tissue is identified in the slides. ResultsWhile we do not control any medical centers protocols for making or storing slides, we developed the iQC tool to hold all medical centers and datasets to the same objective standard of quality. We validate this standard across five Veterans Affairs Medical Centers (VAMCs) and the Automated Gleason Grading Challenge (AGGC) 2022 public dataset. For our surgical procedure prediction task, we report an Area Under Receiver Operating Characteristic (AUROC) of 0.9966-1.000 at the VAMCs that consistently produce high quality data and AUROC of 0.9824 for the AGGC dataset. In contrast, we report an AUROC of 0.7115 at the VAMC that consistently produced poor quality data. An attending pathologist determined poor data quality was likely driven by faded histopathology stains and protocol differences among VAMCs. Corroborating this, iQCs novel stain strength statistic finds this institution has significantly weaker stains (p < 2.2 x 10-16, two-tailed Wilcoxon rank-sum test) than the VAMC that contributed the most slides, and this stain strength difference is a large effect (Cohens d = 1.208). In addition to accurately detecting the distribution of tissue in slides, we find iQC recommends only 2 of 3736 VAMC slides (0.005%) be reviewed for inadequate tissue. With its default configuration file, HistoQC excluded 89.9% of VAMC slides because tissue was not detected in these slides. With our customized configuration file for HistoQC, we reduced this to 16.7% of VAMC slides. Strikingly, the default configuration of HistoQC included 94.0% of the 1172 prostate cancer slides from The Cancer Genome Atlas (TCGA), which may suggest HistoQC defaults were calibrated against TCGA data but this calibration did not generalize well to non-TCGA datasets. For VAMC and TCGA, we find a negligible to small degree of agreement in the include/exclude status of slides, which may suggest iQC and HistoQC are not equivalent. ConclusionOur surgical procedure prediction AUROC may be a quantitative indicator positively associated with high data quality at a medical center or for a specific dataset. We find iQC accurately identifies tissue in slides and excludes few slides, unless the data are poor quality. To produce high quality data, we recommend producing slides using robotics or other forms of automation whenever possible. We recommend scanning slides digitally before the glass slide has time to develop signs of age, e.g faded stains and acrylamide bubbles. We recommend using high-quality reagents to stain and mount slides, which may slow aging. We recommend protecting stored slides from ultraviolet light, from humidity, and from changes in temperature. To our knowledge, iQC is the first automated system in computational pathology that validates data quality against objective evidence, e.g. surgical procedure data available in the EHR or LIMS, which requires zero efforts or annotations from anatomic pathologists. Please see https://github.com/schaumba/iqc and https://doi.org/10.17605/OSF.IO/AVD3Z for instructions and updates.

4
PRECISE: Benchmarking digital pathology with expert-annotated contiguous IHC-H&E serial prostate sections

Calapaqui Teran, A. K.; Gonzalez Bernad, A. A.; Cobo Cano, M.; Sanchez Magdaleno, L.; Marcos Gonzalez, S.; Delgado Bolton, R. C.; Moustafa Calvo, J.; Gomez Roman, J. J.; Lara, L.

2026-07-22 pathology 10.64898/2026.07.21.26358559 medRxiv
Top 0.1%
52.0%
Show abstract

We present PRECISE (PRostate Expert-annotated Contiguous IHC-H\&E Serial sEctions), a hybrid histopathology dataset of paired hematoxylin and eosin (H\&E) and immunohistochemistry (IHC) whole-slide images (WSIs), comprising 37 prostate core needle biopsies from 25 patients, each with matched H\&E and CKAPM+racemase staining. To the best of our knowledge, this is the first publicly available dataset offering spatially harmonized, pixel-level expert annotations across both staining modalities in prostate biopsy WSIs - directly mirroring the two-stage (H\&E-then-IHC) clinical diagnostic workflow used to resolve morphological uncertainty, restricted to cases in which that workflow reached diagnostic consensus. The dataset contains 24,387 annotations spanning seven diagnostically critical classes: malignant glands, benign glands, stromal tissue, intraductal carcinoma (IDC-P), high-grade prostatic intraepithelial neoplasia (HGPIN), atypical intraductal proliferation (AIP), and tissue artifacts. Unlike existing resources, which focus on binary tumor classification or lack IHC pairing, this dataset captures the full morphological spectrum encountered in routine prostate pathology, including rare precursor lesions and confounding entities underrepresented in current benchmarks. Annotations were validated through a structured three-stage consensus by two expert uropathologists, with IHC serving as biological ground truth for boundary definition. PRECISE is designed as a robust benchmark for multimodal semantic segmentation and self-supervised learning, and is openly released to promote reproducible research and accelerate AI-assisted diagnosis in prostate cancer.

5
Cancer-Tissue Fraction as a Scanner-Robust Triage Signal for Automated Gleason Grading of Prostate Biopsies: External Validation Across a Middle Eastern Cohort

Ebbert, J. L.; Perry, A.; Szymanski, J.; Della Corte, D.

2026-07-29 pathology 10.64898/2026.07.28.26359127 medRxiv
Top 0.1%
43.1%
Show abstract

Background: Deep-learning systems for Gleason grading are developed almost entirely on high-end clinical scanners and on cohorts from a small number of Western institutions, yet deployment increasingly involves other devices and other populations. These two distribution shifts, device and population, are rarely tested together on the same physical slides. The PAR dataset, from Erbil, Iraq, digitizes each biopsy on three scanners and provides three distinct pathologist grades, so it permits both tests at once on a Middle Eastern cohort. A concurrent study by the dataset originators validated a task-specific model and two foundation models on PAR; we complement it by testing an inde-pendently developed detect-then-grade pipeline and by separating scanner effects on detection from scanner effects on grading. Methods: We applied one fixed model de-veloped on North American and European material to all 1017 whole-slide images (339 slides from 185 patients, three scanners; 49.6% clinically significant cancer) with no scanner-specific or population-specific tuning. We measured cancer detection (area under the ROC curve of the predicted cancer-tissue fraction), all-slide ISUP agreement of the deployed detect-then-grade pipeline (quadratic-weighted kappa, QWK), and grading agreement on pathologist-confirmed cancers, at the slide level and, because a case carries up to two slides, at the patient level. The reference reader was S.A.; thresholds and operating points were cross-validated leave-one-out; scanners were compared by paired within-biopsy bootstrap and confidence intervals confirmed by patient-cluster bootstrap. Results: Detection was statistically equivalent across scanners (AUC 0.987 to 0.991; paired differences at most 0.003) and transferred to this non-Western cohort with no per-population tuning. At a 95% sensitivity operating point the deployed pipeline reached cross-validated all-slide QWK of 0.86, 0.81, and 0.86 (Grundium, Hamamatsu, Leica), matching the inter-pathologist ceiling of 0.81, against 0.23 to 0.62 for the ungated model. Grading of confirmed cancers was scanner dependent: the compact Grundium (0.63) did not differ from the clinical Leica (0.67; paired difference 0.04, 95% CI -0.03 to 0.11), while both exceeded Hamamatsu (0.44). Results held at the patient level, with grading somewhat lower for two scanners; the two slides of a case disagreed in grade in 43% of cases, and patient clustering did not widen the intervals. Conclusions: Can-cer-tissue fraction is a triage signal robust across scanner and transferable to an un-derrepresented population for detection, while grading is the scanner-sensitive step. Prostate grading models should be deployed as a detect-then-grade pipeline, with grading validated per device and confirmed on the local population.

6
Large-Scale Validation Study of an Improved Semi-Autonomous Urine Cytology Assessment Tool: AutoParis-X

Levy, J.; Chan, N.; Marotti, J.; Kerr, D.; Gutmann, E.; Glass, R.; Dodge, C.; Suriawinata, A.; Christensen, B.; Liu, X.; Vaickus, L.

2023-03-02 pathology 10.1101/2023.03.01.23286639 medRxiv
Top 0.1%
40.7%
Show abstract

Adopting a computational approach for the assessment of urine cytology specimens has the potential to improve the efficiency, accuracy and reliability of bladder cancer screening, which has heretofore relied on semi-subjective manual assessment methods. As rigorous, quantitative criteria and guidelines have been introduced for improving screening practices, e.g., The Paris System for Reporting Urinary Cytology (TPS), algorithms to emulate semi-autonomous diagnostic decision-making have lagged behind, in part due to the complex and nuanced nature of urine cytology reporting. In this study, we report on a deep learning tool, AutoParis-X, which can facilitate rapid semi-autonomous examination of urine cytology specimens. Through a large-scale retrospective validation study, results indicate that AutoParis-X can accurately determine urothelial cell atypia and aggregate a wide-variety of cell and cluster-related information across a slide to yield an Atypia Burden Score (ABS) that correlates closely with overall specimen atypia, predictive of TPS diagnostic categories. Importantly, this approach accounts for challenges associated with assessment of overlapping cell cluster borders, which improved the ability to predict specimen atypia and accurately estimate the nuclear-to-cytoplasm (NC) ratio for cells in these clusters. We developed an interactive web application that is publicly available and open-source, which features a simple, easy-to-use display for examining urine cytology whole-slide images (WSI) and determining the atypia level of specific cells, flagging the most abnormal cells for pathologist review. The accuracy of AutoParis-X (and other semi-automated digital pathology systems) indicates that these technologies are approaching clinical readiness and necessitates full evaluation of these algorithms via head-to-head clinical trials.

7
Using Attention-based Deep Learning to Predict ERG:TMPRSS2 Fusion Status in Prostate Cancer from Whole Slide Images

Omar, M.; Xu, Z.; Rand, S. B.; Mohammad, M.; Salles, D. C.; Schaeffer, E. M.; Robinson, B. D.; Lotan, T. L.; Loda, M.; Marchionni, L.

2022-11-20 pathology 10.1101/2022.11.18.517111 medRxiv
Top 0.1%
39.9%
Show abstract

Prostate cancer (PCa) is associated with several genetic alterations which play an important role in the disease heterogeneity and clinical outcome including gene fusion between TMPRSS2 and members of the ETS family of transcription factors specially ERG. The expanding wealth of pathology whole slide images (WSIs) and the increasing adoption of deep learning (DL) approaches offer a unique opportunity for pathologists to streamline the detection of ERG:TMPRSS2 fusion status. Here, we used two large cohorts of digitized H&E-stained slides from radical prostatectomy specimens to train and evaluate a DL system capable of detecting the ERG fusion status and also detecting tissue regions of high diagnostic and prognostic relevance. Slides from the PCa TCGA dataset were split into training (n=318), validation (n=59), and testing sets (n=59) with the training and validation sets being used for training the model and optimizing its hyperparameters, respectively while the testing set was used for evaluating the performance. Additionally, we used an internal testing cohort consisting of 314 WSIs for independent assessment of the models performance. The ERG prediction model achieved an Area Under the Receiver Operating Characteristic curve (AUC) of 0.72 and 0.73 in the TCGA testing set and the internal testing cohort, respectively. In addition to slide-level classification, we also identified highly attended patches for the cases predicted as either ERG-positive or negative which had distinct morphological features associated with ERG status. We subsequently characterized the cellular composition of these patches using HoVer-Net model trained on the PanNuke dataset to segment and classify the nuclei into five main categories. Notably, a high ratio of neoplastic cells in the highly-attended regions was significantly associated with shorter overall and progression-free survival while high ratios of immune, stromal and stromal to neoplastic cells were all associated with longer overall and metastases-free survival. Our work highlights the utility of deploying deep learning systems on digitized histopathology slides to predict key molecular alteration in cancer together with their associated morphological features which would streamline the diagnostic process.

8
Homologous recombination deficiency prediction from whole slide images using label refinement and foundation-model benchmarking in ovarian cancer

Shah, N. A.; Sarwar, M.; Ullah, E.

2026-06-30 pathology 10.64898/2026.06.25.734452 medRxiv
Top 0.1%
39.8%
Show abstract

Background: Homologous recombination deficiency (HRD) is clinically imperative in high-grade serous ovarian carcinoma (HGSOC), particularly because of its association with platinum sensitivity and benefit from poly(ADP-ribose) polymerase inhibitor (PARPi) therapy. However, public datasets rarely contain a complete combination of diagnostic haematoxylin and eosin (H&E) whole-slide images (WSIs), validated clinical HRD assay results, genomic scar scores, BRCA1 promoter methylation data, and treatment-response outcomes. This creates a major barrier for computational pathology studies seeking to develop clinically interpretable models of HRD or PARPi response from routine histology. Objective: We performed an exploratory, leakage-controlled computational pathology benchmarking study to evaluate whether H&E WSIs from TCGA-OV contain a measurable morphology-linked signal associated with research-grade molecular HRD labels, and whether label refinement and pathology foundation-model embeddings alter predictive performance. Methods: We assembled a frozen-primary TCGA-OV WSI cohort comprising 717 tissue-section/biospecimen slides from 316 patients. Diagnostic FFPE DX slides were excluded from model selection because of complete patient overlap with the frozen-primary cohort. Two HRD labels were evaluated: an initial mutation-only molecular label based on BRCA/HR-gene mutation evidence, and a refined methylation-enhanced molecular label that additionally incorporated BRCA1 promoter methylation. Feature extraction was performed using ResNet50, UNI, CONCH, Virchow2, Phikon-v2, and UNI2-h encoders. Patient-level attention-based multiple instance learning (ABMIL) was used with patient-as-bag modelling. Evaluation used patient-level grouped 5-fold x 5-repeat stratified cross-validation, with 25 folds total, bootstrap confidence intervals, and patient-level leakage control. Results: The initial mutation-only label classified 78 patients as positive and 238 as negative. The refined methylation-enhanced label recovered 33 additional positives, resulting in 111 positive and 205 negative patients. Patient-level ABMIL using UNI2-h features achieved the strongest performance for the refined label, with AUROC 0.634 (95% CI 0.571-0.698), AUPRC 0.468 (95% CI 0.390-0.562), balanced accuracy 0.597, sensitivity 0.532, specificity 0.663, F1 score 0.494, and Brier score 0.233. The calibrated threshold was 0.512, yielding TN=136, FP=69, FN=52, and TP=59. Comparative models showed lower discrimination, including UNI2-h with the initial label (AUROC 0.628), Phikon-v2 refined (0.582), Virchow2 refined (0.582), CONCH initial (0.587), ResNet50 refined (0.570), and clinical baselines (AUROC 0.54-0.57). Conclusions: TCGA-OV H&E WSIs contain a modest but reproducible morphology-linked signal associated with research-grade molecular HRD status. However, the AUROC around 0.63, absence of clinical HRD assay labels, lack of genomic scar endpoints in the implemented workflow, and absence of PARPi/platinum response targets prevent clinical interpretation. This study should be interpreted as a proof-of-concept benchmarking framework and methodological foundation for future H&E-based predictive modelling in clinically curated PARPi response cohorts.

9
Evaluating Large Language Models in Interpreting Cervical Cytology

Geetha, S. D.

2025-11-06 pathology 10.1101/2025.11.04.25339501 medRxiv
Top 0.1%
35.4%
Show abstract

BackgroundLarge language models (LLMs) have shown promise in medical imaging, but their utility in cytology remains underexplored. This study evaluates GPT-5 and Gemini 2.5 Pro for Pap smear interpretation. MethodsDigital cervical Pap smear images of 100 cases were obtained from the Hologic Education Site, with Hologic diagnoses considered the gold standard. Representative images were uploaded into GPT-5 and Gemini 2.5 Pro and prompted to provide a diagnosis based on the Third Edition of the Bethesda System for Reporting Cervical Cytopathology. Cases with infectious organisms were assessed using additional images. Concordance was evaluated at exact diagnosis and clinical management groupings, wherein diagnoses with similar management implications were grouped. Sensitivity and specificity for abnormal cytology were also calculated. ResultsConcordance of both LLMs for exact diagnostic matches were comparable (GPT-5: 47%, Gemini: 48%) and increased to 66% for clinical management grouping. GPT-5 performed best for low-grade squamous intraepithelial lesions (75%), whereas Gemini 2.5 Pro showed the highest concordance in the high-grade squamous intraepithelial lesion (HSIL) category (82%), although this was largely attributable to its strong tendency to overcall cases as HSIL. Sensitivity for detecting abnormal cytology was 74% for GPT-5 and 84% for Gemini, with specificity of 74% and 71%, respectively. GPT-5 better identified glandular lesions, while Gemini detected organisms more accurately (71% vs. 20%). ConclusionsCurrent LLMs demonstrate moderate ability to identify cytologic abnormalities but are not yet reliable for independent Pap smear interpretation. Targeted fine-tuning, prompt optimization, and cytology-specific training could enhance their utility as adjunctive tools in cytology workflows.

10
Unsupervised Tissue Concepts for Explainable Sarcoma Subtype Prediction from H&E

Bisson, T.; Ingram, D.; Singh, S.; Li, A.; Flynn, S.; Wang, W.-L.; Kim, A. E.; Bridge, C. P.; Demicco, E. G.; Sorrentino, A.; Jiang, S.; Hung, Y. P.; Lazar, A. J.; Iafrate, A. J.

2026-05-20 pathology 10.64898/2026.05.15.26353333 medRxiv
Top 0.1%
34.8%
Show abstract

Soft tissue sarcomas are a rare, heterogeneous group of tumors whose diagnosis remains challenging because of overlapping morphology and limited access to sarcoma-specialized pathologists. Although pathology foundation models have shown promise in computational pathology, their clinical translation remains limited by insufficient interpretability, particularly in diagnostically complex settings such as sarcoma diagnosis. Here, we developed and evaluated an H&E-based AI framework for sarcoma subtype classification that focused on explanability. Using the CONCH v1.5 foundation model, we computed embeddings from a tissue microarray cohort of 2,545 cases spanning 19 sarcoma subtypes and trained an attention-based multiple-instance learning model that achieved a balanced accuracy of 77.38% (SD 1.88). To move explainability beyond attention-based localization, we trained a sparse autoencoder on patch-level embeddings to learn 768 recurring visual concepts. 90 high-activation concepts were reviewed by three senior pathologists and curated into morphologically meaningful and non-meaningful categories, yielding a semantic dictionary of 41 diagnostically relevant tissue concepts. We then trained a linear attention-based model on the 768-concept vectors, which retained much of the performance of the raw embedding-based ABMIL model, achieving a balanced accuracy of 73.74% (SD 1.30). When restricting the linear model to pathologist-curated morphologic concepts only, balanced accuracy further decreased to 67.04% (SD 1.27), suggesting that the residual performance gain in the full concept model was driven by inconsistent, technical, or diagnostically irrelevant concepts. Concept-level explanations of the curated linear attention-based model aligned with known sarcoma morphology, including lipogenic, myxoid, spindle-cell, pleomorphic, vascular, small round blue cell, and matrix-forming patterns, and reproduced patterns of diagnostic overlap observed in human sarcoma pathology. Together, these results show that H&E-based foundation-model representations capture meaningful diagnostic structure within the known limitations of H&E in sarcoma diagnostics, but that their clinical value depends on whether this structure can be made interpretable to pathologists. Sparse autoencoder-derived concepts can address this critical gap by converting embedding-level signal into recurring morphologic patterns that pathologists can review and name, providing the foundation to link these patterns to subtype predictions. In doing so, this approach turns concept discovery into a practical form of diagnostic explanation, while also revealing where model performance is supported by recognizable histopathology and where it relies on diagnostically irrelevant or inconsistent visual patterns.

11
Comparison of machine learning to deep learning for automated annotation of Gleason patterns in whole mount prostate cancer histology.

Duenweg, S. R.; Brehler, M.; Bobholz, S. A.; Lowman, A. K.; Winiarz, A.; Kyereme, F.; Nencka, A.; Iczkowski, K. A.; LaViolette, P.

2022-11-14 pathology 10.1101/2022.11.10.516007 medRxiv
Top 0.1%
34.5%
Show abstract

BackgroundOne in eight men will be affected by prostate cancer (PCa) in their lives. While the current clinical standard prognostic marker for PCa is the Gleason score, it is subject to interreviewer variability. This study compares two machine learning methods for discriminating between high- and low-grade PCa on histology from 47 PCa patients. MethodsDigitized slides were annotated by a GU fellowship-trained pathologist. High-resolution tiles were extracted from annotated and unlabeled tissue. Glands were segmented and pathomic features were calculated and averaged across each patient. Patients were separated into a training set of 31 patients (Cohort A, n=9345 tiles) and a testing cohort of 16 patients (Cohort B, n=4375 tiles). Tiles from Cohort A were used to train a compact classification ensemble model and a ResNet model to discriminate tumor and were compared to pathologist annotations. ResultsThe ensemble and ResNet models had overall accuracies of 89% and 88%, respectively. The ResNet model was additionally able to differentiate Gleason patterns on data from Cohort B while the ensemble model was not. ConclusionsOur results suggest that quantitative pathomic features calculated from PCa histology can distinguish regions of cancer; how-ever, texture features captured by deep learning frameworks better differentiate unique Gleason patterns.

12
LymphoML: An interpretable artificial intelligence-based method identifies morphologic features that correlate with lymphoma subtype

Shankar, V.; Yang, X.; Krishna, V.; Tan, B.; Silva, O.; Rojansky, R.; Ng, A.; Valvert, F.; Briercheck, E.; Weinstock, D.; Natkunam, Y.; Fernandez-Pol, S.; Rajpurkar, P.

2023-03-17 pathology 10.1101/2023.03.14.23287143 medRxiv
Top 0.1%
34.5%
Show abstract

Lymphomas vary in terms of clinical behavior, morphology, and response to therapies and thus accurate classification is essential for appropriate management of patients. In this study, using a set of 670 cases of lymphoma obtained from a center in Guatemala City, we propose an interpretable machine learning method, LymphoML, for lymphoma subtyping into eight diagnostic categories. LymphoML sequentially applies steps of (1) object segmentation to extract nuclei, cells, and cytoplasm from hematoxylin and eosin (H&E)-stained tissue microarray (TMA) cores, (2) feature extraction of morphological, textural, and architectural features, and (3) aggregation of per-object features to create patch-level feature vectors for lymphoma classification. LymphoML achieves a diagnostic accuracy of 64.3% (AUROC: 85.9%, specificity: 88.7%, sensitivity: 66.9%) among 8 lymphoma subtypes using only H&E-stained TMA core sections, at a level similar to experienced hematopathologists. We find that the best models set of nuclear and cytoplasmic morphological, textural, and architectural features are most discriminative for diffuse large B-cell lymphoma (F1: 78.7%), classic Hodgkin lymphoma (F1 score: 74.5%), and mantle cell lymphoma (F1: 71.0%). Nuclear shape features provide the highest diagnostic yield, with nuclear texture, cytoplasmic, and architectural features providing smaller gains in accuracy. Finally, combining information from the H&E-based model together with the results of a limited set of immunohistochemical (IHC) stains resulted in a similar diagnostic accuracy (accuracy: 85.3%, AUROC: 95.7%, sensitivity: 84.5%, specificity: 93.5%) as with a much larger set of IHC stains (accuracy: 86.1%, AUROC: 96.7%, specificity: 93.2%, sensitivity: 86.0%). Our work suggests a potential way to incorporate machine learning tools into clinical practice to reduce the number of expensive IHC stains while achieving a similar level of diagnostic accuracy.

13
Histological triage of early-stage mycosis fungoides using a weakly supervised deep learning-based model: a multicentre, external validation, and clinical utility study

Doeleman, T.; Brussee, S.; Valkema, P.; Kempf, W.; Vermeer, M.; Kers, J.; Wynaendts, L.; Kerckhoffs, K.; de Jonge, M.; Nguyen, A.; Peters, E.; Wobser, M.; Rauert-Wunderlich, H.; Rosenwald, A.; Stadler, R.; Jansen, P.; Battistella, M.; Roccuzzo, G.; Quaglino, P.; Schrader, A.

2026-07-28 pathology 10.64898/2026.07.27.26359009 medRxiv
Top 0.1%
31.4%
Show abstract

Background Histological diagnosis of early-stage mycosis fungoides (MF) is hindered by profound overlap with benign inflammatory dermatoses (BIDs), leading to diagnostic delays and extensive ancillary testing. We developed MIMIC (Multiple Instance-learning for Identification of Mycosis fungoides In Cutaneous biopsies), a weakly supervised deep learning model designed as a triage tool at initial H&E whole slide image (WSI) review to distinguish classic patch and plaque stage MF from BIDs. We externally validated the model and evaluated its clinical utility. Methods In this retrospective multicentre study, we trained a base model using weakly supervised attention based multiple instance learning on 3,339 WSIs from two Dutch centres. Crucially, all MF training labels were derived from a deeply phenotyped national cohort featuring strict multidisciplinary expert panel consensus diagnoses (the clinical gold standard). Transportability was evaluated on 371 WSIs from four independent European centres. A blinded reader study on 171 WSIs compared morphology only performance of MIMIC with 11 (dermato-)pathologists. We then retrained an updated model on all retrospective multicentre data and assessed clinical utility in a strictly held out, consecutive Utrecht cohort (2022-2023; 486 accessions, 863 WSIs). Primary analysis focused on classic MF versus BIDs (453 accessions). Decision curve analysis, using Platt scaled probabilities to correct for spectrum bias, evaluated net benefit at a prespecified, safety oriented threshold of 0.04. Findings The base model showed good multicentre transportability (mean centre specific AUROC 0.91; pooled AUROC 0.84). In the reader study, MIMIC achieved an AUROC of 0.87, exceeding the mean pathologist AUROC (0.79) and the best individual reader (0.83). In the consecutive MF versus BID cohort, the updated model achieved an AUROC of 0.87 (95% CI 0.81-0.92). At the 0.04 threshold, sensitivity was 97.8% (44/45 MF cases) and specificity 50.2%, reducing unnecessary ancillary workups by 39.9 per 100 screening cases versus a test all strategy. Interpretation By identifying nearly half of BIDs as low risk while preserving near complete sensitivity for classic early stage MF in a European digital pathology workflow, this unimodal H&E approach offers a scalable digital solution to reduce defensive ancillary testing and accelerate the diagnostic journey for patients with MF. Further validation is needed in non European centres and in populations with darker skin phototypes.

14
The future of computational pathology: expectations regarding the anticipated role of artificial intelligence in pathology by 2030

Berbis, M. A.; Berbis, M. A.; McClintock, D. S.; Bychkov, A.; Cheng, J. Y.; Delahunt, B.; Egevad, L.; Eloy, C.; Farris, A. B.; Fraggetta, F.; Garcia del Moral, R.; Hartman, D. J.; Herrmann, M. D.; Hollemans, E.; Iczkowski, K. A.; Karsan, A.; Kriegsmann, M.; Lennerz, J. K.; Pantanowitz, L.; Salama, M. E.; Sinard, J.; Tuthill, M.; Van der Laak, J.; Williams, B.; Casado-Sanchez, C.; Casado-Sanchez, C.; Sanchez-Turrion, V.; Sanchez-Turrion, V.; Luna, A.; Aneiros-Fernandez, J.; Aneiros-Fernandez, J.; Shen, J.

2022-09-04 pathology 10.1101/2022.09.02.22279476 medRxiv
Top 0.1%
31.2%
Show abstract

BackgroundArtificial intelligence (AI) is rapidly fueling a fundamental transformation in the practice of pathology. However, AIs clinical integration remains challenging, with no AI algorithms to date enjoying routine adoption within typical anatomic pathology (AP) laboratories. This survey gathered current expert perspectives and expectations regarding the role of AI in AP from those with first-hand computational pathology and AI experience. MethodsPerspectives were solicited using the Delphi method from 24 subject matter experts between December 2020 and February 2021 regarding the anticipated role of AI in pathology by the year 2030. The study consisted of three consecutive rounds: 1) an open-ended, free response questionnaire generating a list of survey items; 2) a Likert-scale survey scored by experts and analyzed for consensus; and 3) a repeat survey of items not reaching consensus to obtain further expert consensus. FindingsConsensus opinions were reached on 141 of 180 survey items (78.3%). Experts agreed that AI would be routinely and impactfully used within AP laboratory and pathologist clinical workflows by 2030. High consensus was reached on 100 items across nine categories encompassing the impact of AI on (1) pathology key performance indicators (KPIs) and (2) the pathology workforce and specific tasks performed by (3) pathologists and (4) AP lab technicians, as well as (5) specific AI applications and their likelihood of routine use by 2030, (6) AIs role in integrated diagnostics, (7) pathology tasks likely to be fully automated using AI, and (8) regulatory/legal and (9) ethical aspects of AI integration in pathology. InterpretationThis is the first systematic consensus study detailing the expected short/mid-term impact of AI on pathology practice. These findings provide timely and relevant information regarding future care delivery in pathology and raise key practical, ethical, and legal challenges that must be addressed prior to AIs successful clinical implementation. FundingThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.

15
NKG2C Improves Diagnostic Specificity of NK Cell Receptor Restriction by Identifying Non-Neoplastic Adaptive NK Cell Clones

Wilk, A. J.; Gitana, G.; Oak, J.

2025-11-22 pathology 10.1101/2025.11.18.25340429 medRxiv
Top 0.1%
30.8%
Show abstract

Natural killer (NK) cell neoplasms are a diverse group of entities with often nonspecific clinical presentations, making immunophenotyping essential for diagnosis. Immunophenotyping by flow cytometry can identify clonal NK cell populations by detecting restricted expression patterns of NK cell receptors such as killer cell immunoglobulin-like receptors (KIRs). However, reactive NK cells may also demonstrate KIR restriction through expansion of self-KIR-expressing NK cells, leading to identification of NK clones of uncertain significance (NK-CUS). A well-described reactive NK subset, termed "adaptive" NK cells, arises in response to cytomegalovirus (CMV) infection or reactivation, often appears KIR-restricted, and is defined by coexpression of CD57 and the activating receptor NKG2C. Because CMV reactivation is common among patients undergoing evaluation for hematolymphoid malignancy, we hypothesized that NK-CUS may frequently correspond to this non-neoplastic adaptive NK cell subset. Here, we describe a flow cytometry panel for immunophenotypic characterization of cytotoxic lymphocytes that includes NKG2C, enabling detection of non-neoplastic adaptive NK cells. We show that NK-CUS frequently represent reactive NKG2C+ adaptive NK cells. We describe several cases that meet diagnostic criteria for NK-large granular lymphocytic leukemia (NK-LGLL) and demonstrate that the NK cell clones are non-neoplastic NKG2C+ adaptive NK cells arising in the setting of CMV viremia. Further, we show that NKG2C expression is uncommon by cytotoxic lymphocyte malignancies with recurrent molecular or cytogenetic abnormalities. Collectively, we demonstrate that NKG2C has a high specificity for reactive NK cell populations, and its inclusion in NK cell immunophenotyping panels is a useful strategy to more reliably distinguish between neoplastic and reactive NK cell populations.

16
Pathology's Last Exam: Stress-Testing Diagnostic Reasoning and Safety in Large Language Models

Reitsam, N. G.; Gustav, M.; Jesinghaus, M.; Maerkl, B.; Foersch, S.; Kather, J. N.

2025-12-15 pathology 10.64898/2025.12.11.25342081 medRxiv
Top 0.1%
27.5%
Show abstract

Large language models (LLMs) are evolving into diagnostic co-pilots, yet current benchmarks fail to test the integrated, stepwise reasoning required in diagnostic pathology. Here, we present Pathologys Last Exam (PLE), a curated, highly detailed, text-based benchmark of 100 complex cases spanning organ systems, enriched for rare/challenging entities, plus 20 adversarial cases designed to stress-test model safety. Each case provides structured blocks (Primary, Clinical, Histopathology, IHC/Special Stains, Molecular Pathology) with stepwise information release mirroring real sign-out. We evaluated five LLMs (one proprietary, four open-source) across different stages. While the best model (GPT-5) achieved 70% accuracy on full evidence, performance on safety tests was alarming. Models frequently failed to detect biological contradictions, confidently diagnosing nonsensical "mix-up" cases rather than refusing them. This reveals a critical safety gap: high diagnostic capability is currently coupled with a dangerous inability to recognize impossible clinical scenarios. PLE provides a framework to measure and mitigate these risks before clinical deployment, as well as a foundation for developing multimodal evaluation protocols that can be extended to vision-language models and autonomous diagnostic agents in the future.

17
Dendrite: A Structured, Accessible, and Queryable Pathology Search Database for Streamlined Experiment Planning

Lu, Y.; Hamilton, R.; Greenburg, J.; Srinivasan, G.; Shah, P.; Preum, S.; Pettus, J.; Vaickus, L.; Levy, J.

2023-09-10 pathology 10.1101/2023.09.09.23295302 medRxiv
Top 0.1%
27.4%
Show abstract

Pathology reports contain vital information, yet a significant portion of this data remains underutilized in electronic medical record systems due to the unstructured and varied nature of reporting. Although synoptic reporting has introduced reporting standards, the majority of pathology text remains free-form, necessitating additional processing to enable accessibility for research and clinical applications. This paper presents Dendrite, a web application designed to enhance pathology research by providing intelligent search capabilities and streamlining the creation of study cohorts. Leveraging expert knowledge and natural language processing algorithms, Dendrite converts free-form pathology reports into structured formats, facilitating easier querying and analysis. Using a custom Python script, Dendrite organizes pathology report data, enabling record linkages, text searches, and structured drop-down menus for information filtering and integration. A companion web application enables data exploration and export, showcasing its potential for further analysis and research. Dendrite, derived from existing laboratory information systems, outperforms existing implementations in terms of speed, responsiveness, and flexibility. With its efficient search functionality and support for clinical research and quality improvement efforts in the pathology field, Dendrite proves to be a valuable tool for pathologists. Future enhancements encompass user management integration, integration of natural language processing and machine learning to enhance structured reporting capabilities and seamless integration of Dendrite with the vast repository of genomics and imaging data.

18
Examining Longitudinal Markers of Bladder Cancer Recurrence Through a Semi-Autonomous Machine Learning System for Quantifying Specimen Atypia from Urine Cytology

Levy, J.; Chan, N.; Marotti, J.; Rodrigues, N.; Ismail, A. A.; Kerr, D.; Gutmann, E.; Glass, R.; Dodge, C.; Suriawinata, A.; Christensen, B.; Liu, X.; Vaickus, L.

2023-03-05 pathology 10.1101/2023.03.02.23286716 medRxiv
Top 0.1%
27.3%
Show abstract

Urine cytology (UC) is generally considered the primary approach for screening for recurrence of bladder cancer. However, it is currently unclear how best to use cytological exams themselves for the assessment and early detection of recurrence, beyond identifying a positive finding which requires more invasive methods to confirm recurrence and decide on therapeutic options. As screening programs are frequent, and can be burdensome, finding quantitative means to reduce this burden for patients, cytopathologists and urologists is an important endeavor and can improve both the efficiency and reliability of findings. Additionally, identifying ways to risk-stratify patients is crucial for improving quality of life while reducing the risk of future recurrence or progression of the cancer. In this study, we leveraged a computational machine learning tool, AutoParis-X, to extract imaging features from UC exams longitudinally to study the predictive potential of urine cytology for assessing recurrence risk. This study examined how the significance of imaging predictors changes over time before and after surgery to determine which predictors and time periods are most relevant for assessing recurrence risk. Results indicate that imaging predictors extracted using AutoParis-X can predict recurrence as well or better than traditional cytological / histological assessments alone and that the predictiveness of these features is variable across time, with key differences in overall specimen atypia identified immediately before tumor recurrence. Further research will clarify how computational methods can be effectively utilized in high volume screening programs to improve recurrence detection and complement traditional modes of assessment.

19
Artificial Intelligence Helps to Predict Recurrence and Mortality for Prostate Cancer using Histology Images

Eminaga, O.; Saad, F.; Tian, Z.; Wolffgang, U.; Karakiewicz, P. I.; Ouellet, V.; Azzi, F.; Spieker, T.; Helmke, B. M.; Graefen, M.; Jiang, X.; Xing, L.; Witt, J. H.; Trudel, D.; Leyh-Bannurah, S.-R.

2023-08-20 pathology 10.1101/2023.07.27.550781 medRxiv
Top 0.1%
27.2%
Show abstract

Besides grading, deep learning could improve expert consensus to predict prostate cancer (PCa) recurrence. We developed a novel PCa recurrence prediction system based on artificial intelligence (AI). We validated it using multi-institutional and international datasets comprising 2,647 PCa patients with at least a 10-year follow-up. Survival analyses were performed and goodness-of-fit of multivariate models was evaluated using partial likelihood ratio tests, Akaikes test, or Bayesian information criteria to determine the superiority of our system over existing grading systems. Comprehensive survival analyses demonstrated the effectiveness of our AI- system in categorizing PCa into four distinct risk groups. The system was independent and superior to the existing five grade groups for malignancies. A high consensus level was observed among five blinded genitourinary pathology experts in ranking images according to our prediction system. Therefore, AI may help develop an accurate and clinically interpretable PCa recurrence prediction system, facilitating informed decision-making for PCa patients.

20
Impact of variation in tissue staining and scanning devices on performance of pan-cancer AI models: a study of sarcoma and their mimics

Chai, B.; Chen, J.; Cool, P.; Oumlil, F.; Tollitt, A.; Steiner, D. F.; Chakraborti, T.; Flanagan, A. M.

2025-08-22 pathology 10.1101/2025.08.18.670932 medRxiv
Top 0.1%
26.4%
Show abstract

Histopathological analysis is considered the gold standard for the diagnosis and prognostication of cancer. Recent advances in AI, driven by large-scale digitisation and pan-cancer foundation models, are opening new opportunities for clinical integration. However, it remains unclear how robust these foundation models are to real-world sources of variability, particularly in H&E staining and scanning protocols. In this study, we use soft tissue tumours, a rare and morphologically diverse tumour type, as a challenging test case to systematically investigate the colour-related robustness and generalisability of seven AI models. Controlled staining and scanning experiments were utilised to assess model performance across diverse real-world data sources. Foundation models, particularly UNI-v2, Virchow and TITAN, demonstrated encouraging robustness to staining and scanning variation, particularly when a small number of stain-varied slides were included in the training loop, highlighting their potential as adaptable and data-efficient tools for real-world digital pathology workflows.